SemanticScuttle - klotz.me » Tags: evaluation+large language models

Tags: evaluation* + large language models*

0 bookmark(s) - Sort by: Date ↓ / Title /

How to Implement the LLM Arena-as-a-Judge Approach to Evaluate Large Language Model Outputs

This tutorial explores implementing the LLM Arena-as-a-Judge approach to evaluate large language model outputs using head-to-head comparisons. It demonstrates using OpenAI’s GPT-4.1 and Gemini 2.5 Pro, judged by GPT-5, in a customer support scenario.

2025-08-26 Tags: llm, arena-as-a-judge, evaluation, openai, gpt-4, gemini, gpt-5, deepeval, machine learning by klotz

asta-paper-finder

frozen-in-time version of our Paper Finder agent for reproducing evaluation results. This repo contains the code for the standalone Paper Finder agent. PaperFinder is our paper-seeking agent, which is intended to assist in locating sets of papers according to content-based and metadata criteria.

2025-08-26 Tags: paper finder, agent, llm, research papers, evaluation, python by klotz

LLM Evaluation

This GitHub repository directory contains resources for evaluating Large Language Models (LLMs), including a Jupyter Notebook demonstrating how to use LLM Arena as a judge and a Python script for the same purpose. It also includes a README file with instructions on how to view the notebook if it doesn't render correctly on GitHub.

2025-08-26 Tags: llm, evaluation, large language models, llm arena, jupyter notebook, python, ai, github by klotz

MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers

MCP-Universe is a comprehensive benchmark designed to evaluate LLMs in realistic tasks through interaction with real-world MCP servers across 6 core domains and 231 tasks. It highlights the challenges of long-context reasoning, unfamiliar tool spaces, and cross-domain variations in LLM performance.

2025-08-25 Tags: llm, benchmark, mcp, model context protocol, evaluation, agent by klotz

From Prototype to Production: Enhancing LLM Accuracy

This article discusses methods to measure and improve the accuracy of Large Language Model (LLM) applications, focusing on building an SQL Agent where precision is crucial. It covers setting up the environment, creating a prototype, evaluating accuracy, and using techniques like self-reflection and retrieval-augmented generation (RAG) to enhance performance.

2024-12-20 Tags: llm, accuracy, evaluation, sql, agent, rag by klotz

We Ran Over Half a Million Evaluations on Quantized LLMs: Here's What We Found

This article discusses the extensive evaluation of quantized large language models (LLMs) by Neural Magic, finding that quantized LLMs maintain competitive accuracy and efficiency with their full-precision counterparts.

- **Quantization Schemes**: Three different quantization schemes were tested: W8A8-INT, W8A8-FP, and W4A16-INT, each optimized for different hardware and deployment scenarios.
- **Accuracy Recovery**: The quantized models demonstrated high accuracy recovery, often reaching over 99%, across a range of benchmarks, including OpenLLM Leaderboard v1 and v2, Arena-Hard, and HumanEval.
- **Text Similarity**: Text generated by quantized models was found to be highly similar to that generated by full-precision models, maintaining semantic and structural consistency.

2025-02-27 Tags: quantization, llm, evaluation, neural magic by klotz

Metrics to Evaluate a Classification Machine Learning Model

This article explores various metrics used to evaluate the performance of classification machine learning models, including precision, recall, F1-score, accuracy, and alert rate. It explains how these metrics are calculated and provides insights into their application in real-world scenarios, particularly in fraud detection.

2024-08-01 Tags: machine learning, classification, metrics, evaluation, precision, recall, f1-score, accuracy, alert rate, fraud detection, llm by klotz

End-to-end LLM Workflows Guide

This guide demonstrates how to execute end-to-end LLM workflows for developing and productionizing LLMs at scale. It covers data preprocessing, fine-tuning, evaluation, and serving.

2024-06-21 Tags: llm, workflows, data preprocessing, fine-tuning, evaluation, serving, ray, anyscale by klotz

New Trends in LLM Architecture

Discusses the trends in Large Language Models (LLMs) architecture, including the rise of more GPU, more weights, more tokens, energy-efficient implementations, the role of LLM routers, and the need for better evaluation metrics, faster fine-tuning, and self-tuning.

2024-06-01 Tags: llm, machine learning, deep learning, transformers, self-tuning, evaluation by klotz

Langfuse - Open Source LLM Engineering Platform

Langfuse is an open-source LLM engineering platform that offers tracing, prompt management, evaluation, datasets, metrics, and playground for debugging and improving LLM applications. It is backed by several renowned companies and has won multiple awards. Langfuse is built with security in mind, with SOC 2 Type II and ISO 27001 certifications and GDPR compliance.

2024-05-23 Tags: lamgfuse, llm, prompt engineering, evaluation, datasets, metrics, observability by klotz

First / Previous / Next / Last / Page 1 of 0

SemanticScuttle - klotz.me

Tags: evaluation* + large language models*

Linked Tags

Related Tags